The story of a competition entry
Teaching a model to teach
I picked Math & Scientific Reasoning because pragmatic education is something African classrooms need and rarely get enough of: too many students for one teacher, and too little time to slow down on the question any single student actually got wrong.
The competition evaluates an offline model on a $150–$500 CPU-only laptop. The model must therefore balance teaching accuracy, response speed, and memory use. This report follows the experiments that narrowed those three variables to two runtime-dependent finalists.
Where the story stands, for now
Fine-tuning improved both finalists
Qwen3.5 0.8B remains the scalar leader after a 15-point ARC-Easy gain. Qwen2.5 1.5B remains the vector leader after a smaller 3.4-point gain. These are the best measured GGUFs, not yet the packaged submission: both still need the tutor metadata and live-prompt checks, and the final choice still depends on the audit CPU configuration.
Scalar leader, n=500Qwen3.5 0.8B Q4_0
- Scalar total
- 80.3664
- Vector total
- 82.3962
- Scalar generation
- 13.60 tok/s
- Est. scalar RSS
- 691 MiB
- ARC-Easy, n=500
- 70.2%
Vector leader, n=500Qwen2.5 1.5B Q4_K_M
- Vector total
- 84.1387
- Scalar total
- 67.0475
- Vector generation
- 17.44 tok/s
- Est. vector RSS
- 1,706 MiB
- ARC-Easy, n=500
- 77.8%
Selection condition: use the tuned Qwen3.5 candidate for the supplied scalar configuration; use tuned Qwen2.5 only if the audit is confirmed to use the vector configuration. Before submission, apply the same embedded tutor policy to the chosen GGUF and repeat the live-prompt and physical-target checks.
Loaded evidence differs from the report snapshot. The operational appendix shows the configured campaign.
- Scalar candidate size
- 513.0 MB
- Quantization
- Q4_0
- Target
- 8 GB, CPU-only
- Selection basis
- Matched n=500 accuracy, speed, and RSS
The rules we're playing by
A $150–$500 laptop is the whole target
The Africa Deep Tech Challenge wants language models running entirely on-device, on the standard laptops already sitting in African cities, with no cloud behind them. It lists seven tracks: Math & Scientific Reasoning, Healthcare, Agriculture, Creative Writing, Coding Assistants, Corporate/Enterprise, and Autonomous Agents. We took the first one, because a tutor that explains a wrong answer instead of just marking it is closer to what a scarce teacher actually does.
The hardware is exact: an Intel Core i5, 10th to 12th generation, integrated graphics only, Ubuntu 22.04, and a hard 7 GB RAM ceiling. Go over it and the run is disqualified, full stop. The rules are just as narrow on software: llama.cpp only, weights in GGUF. Any open base model is fair game, quantized or fine-tuned however we like, but nothing closed and nothing needing a different runtime.We disable AVX-512 in our own build and check the disassembly for it, because several eligible consumer CPUs don't support it: an illegal instruction there means an automatic zero.
Half the score is accuracy, so a fast model that can't teach still loses. But performance caps at 15 tok/s and efficiency zeroes out at 7 GB, so past those points speed and memory stop paying at all.S_perf = min(TPS/15, 1) × 100 and S_eff = max(0, (7 GB − RSS)/7 GB) × 100. The public challenge page instead defines performance relative to the fastest submission, TPS_max, rather than a fixed 15. We follow the fixed formula, since it's what the executable profiler actually runs, and keep the cohort-relative reading as a separate sensitivity check throughout this page. That formula produces a decision rule we checked every optimisation against for the rest of the project:
Two CPU configurations run through this story, named once so we can stop repeating the acronyms. The scalar configuration is the supplied profiler build, with the wider vector extensions disabled for the affected kernels. The vector configuration is a portable SIMD build with AVX2, FMA, and F16C enabled, on the same two physical cores. From here on, this page just says scalar and vector.
Why we keep five separate evidence categories on this page
Our numbers come from different systems, CPU features, and performance formulas, so we never blend them into one ranking: official profiler results (the participant executable's own runs), profiler-parity estimates (a scalar-only reconstruction with a 45 MiB profiler-root allowance we estimated, not measured), controlled vector proxies (a paired scalar/vector benchmark on the same host), website sensitivity (the cohort-relative formula as a separate check, never averaged in), and plain development results (Mac, Docker, and custom-engine tests that guided decisions but never ranked a submission). A colour and a badge on every figure below show which lane it belongs to.
Before we picked a model
Bigger teaches better. Bigger also costs more.
Before touching quantization, we ran a plain accuracy-by-scale study across the Qwen3.5 family: 0.8B, 2B, and 4B, on four reasoning tasks. The result was the least surprising thing in this whole project, and also the thing every later decision had to argue against.
Separate accuracy evaluations
Candidate measurements
The separate four-task reasoning study recorded Qwen3.5 4B at 2.55 GiB and 73.3, 2B at 1.19 GiB and 64.8, and 0.8B at 0.50 GiB and 51.3. The audit-candidate panel uses separate ARC-Easy proxies. Fine-tuned Qwen3.5 0.8B is 0.48 GiB at 70.2% and leads the scalar total. Fine-tuned Qwen2.5 1.5B is 0.92 GiB at 77.8% and leads the vector total. Earlier candidates remain as comparison points.
4B beat 2B beat 0.8B, cleanly, on every task. Under a formula that pays 0.50 total-score points per accuracy point and only 2.86 per gigabyte saved, that curve makes a strong case for going as big as the RAM ceiling allows. But RSS on Linux tracks file bytes almost one-to-one, so "as big as the ceiling allows" is a number in the hundreds of megabytes, not gigabytes, once quantization is priced in. That tension split our search down two roads at once, plus one number neither road was allowed to cross:
Road one
Go small and dense
Stay in the Qwen3/3.5 family, quantize as hard as the audit kernel would tolerate, and win back accuracy through parameter dedication and behaviour baked into the file rather than raw scale.
Road two
Go big and ternary
Bet that an 8B-parameter model at roughly 2 bits per weight, BitCPM4 in TQ2_0, could out-accuracy anything dense enough to fit, if we pruned and compressed it hard enough to land in the same weight class as the dense candidates.
Every 100 MB above the smallest working file was already worth 1.4 points of efficiency score we'd rather spend on accuracy, so neither road got a free pass on size: road two's ternary bet only earned its keep by landing at about 2.1 GB, not the 8 GB the parameter count alone would suggest. We ran both roads in parallel for weeks. This page tells road one's story in full, because it's the one that shipped; road two still shows up wherever its lessons apply, and its ending is in the vocabulary-pruning section below.
The toolbox
Ten optimization methods
We tested ten ways to change accuracy, file size, RAM, or speed. Each method either narrowed the candidate set or was retained as a measured rejection.
Quantization
Which quantization type the audit binary can actually run fast, not just how many bits per weight.
Parameter deduplication
Dropping the duplicated output head when embedding and head can share one matrix.
Vocabulary pruning
Removing tokens the tutor will never emit, and checking token by token that nothing else changed.
Weight streaming
Paging weights from disk instead of keeping the whole file resident, and finding the boundary the competition draws around it.
Runtime tuning
Threads, KV cache, checkpoints, and two kinds of speculative decoding, before any model comparison could be trusted.
Instruction-set targeting
Discovering that the audit binary's actual CPU features change which quantization type wins, and testing both readings.
Widening the model search
Testing nine more candidates and a specialist fine-tune once the first roster stopped improving.
Baking behaviour into the file
Persona, chat template, and sampling defaults, since judges talk to the bare GGUF with no app layer in front of it.
Fine-tuning
Updating the two finalists on verified math and science questions matched to the profiler's continuation format.
Runtime configuration we couldn't submit
Repacking, mmap policy, thread count: measured changes that live in Muta's runtime, not in the submitted GGUF.
Quantization, deduplication, embedded behaviour, and fine-tuning can change the submitted GGUF. Streaming and runtime tuning cannot. The model search identifies which architecture receives those changes; the scalar/vector comparison determines how each candidate is scored.
Method one
Quantization: the file format is the kernel choice
Quantization means storing each weight in fewer bits than the 16 or 32 it was trained in, trading a little precision for a much smaller file. GGUF supports a whole ladder of these formats, from plain 4-bit rounding (Q4_0) through vector-quantized "k-quants" (Q4_K_M, Q5_K_M) to sub-4-bit "i-quants" and ternary formats. We assumed, like most guides do, that a newer, smarter format like Q4_K_M would be the better choice. On this evaluator, it wasn't.
Scalar profiler configuration
Kernel dispatch by tensor type
On the scalar participant build, Q4_0 tensors use a hand-written SSSE3 kernel and the current Qwen final decodes at 12.63 tokens per second. K-quants, i-quants, and ternary types use scalar generic C implementations; their measured rates span 0.81 to 12.72 tokens per second across the current six-model ledger because model scale also differs. The vector build enables SIMD kernels for these tensor types and changes the quant ordering, without changing any model's measured accuracy.
We read the audit binary's own kernel source rather than trust a guide written for a different CPU, and found only Q4_0 has a hand-written SIMD path on the scalar build the executable profiler runs. Everything else, including the "smarter" k-quants, falls back to generic C. The original control, on the Qwen3 1.7B pair, measured 9.99 tok/s in Q4_0 against 5.30 tok/s in Q4_K_M, purely from that one format choice; the diagram above shows the same pattern holding on the file we eventually shipped, at 12.63 tok/s. That single fact ruled out most of the usual advice and pointed every candidate toward pure Q4_0, at least until instruction-set targeting complicated the picture again, several sections down.
Method two
Cutting the duplicate head
Many transformer architectures give the embedding table and the output head separate weight matrices, even though both are the same shape and often learn near-identical representations. Tying them, letting the head reuse the embedding matrix instead of storing its own copy, is a free win if the architecture supports it: no retraining, no accuracy cost, just a smaller file.
We reproduced our recommended file from its source quant using metadata changes alone, then ran an isolated A/B on the tied-versus-untied question: removing the duplicate saved about 175 MB with identical ARC-Easy accuracy in both conditions, an unambiguous win we took on every candidate that supported it.
The tutor persona, chat template, and sampling defaults go in through a similar scripted pass — though that one only touches metadata, not tensors, so it costs no file size at all. It gets its own full treatment in the behaviour-baking section further down, once the mechanisms that shrink the file are out of the way.
Method three, and the end of road two
Pruning the vocabulary, and letting the ternary bet go
Vocabulary pruning removes tokens a model will never realistically emit from its embedding table, shrinking the largest single matrix in a small model without touching a single weight the tutor actually uses. We built this for road two, the 8B ternary BitCPM4 model, since an English tutor has no use for most of a 73,448-token vocabulary built to cover CJK scripts.
Vocabulary pruning
We cut the vocabulary from 73,448 to 44,416 tokens and padded it to a multiple of 64, which shrank the file by 164 MB. English tokenisation matched across 20,464 checked tokens; perplexity fell from 10.558 to 10.473.
KeptTQ1_0 body test
The body shrank by 340 MB, but generic CPU throughput fell from 3.70 to 2.88 tok/s. The lower bit width bought us nothing on the evaluated kernel's instruction cost.
RejectedHead and embedding quantization
Head and embedding requantization saved at most 48 MB. The largest estimated total-score gain we found was about 0.14 points, too small to matter.
RejectedFactorisation and sparsity tests
The ternary matrices turned out to be full-rank: rank-2048 factorisation error was about 0.80 before quantization. Dense GGUF storage and kernels give us no credit for unstructured zeros either, so sparsity saved nothing.
RejectedOnly the vocabulary prune was worth keeping, and it was the last real win road two got. BitCPM4-8B-TQ2_0 still recorded the highest accuracy of any candidate we ever tested, 88% ARC-Easy, well above every Qwen quantization we ran. But it decoded at just 0.81 tok/s on the scalar audit kernel, since ternary formats fall on the same generic-C path as k-quants. Even after instruction-set targeting later gave it a vector kernel and lifted that to 7.49 tok/s, its total reached only 72.5121, still short of every Qwen variant. The most accurate model we built turned out to be the one we couldn't ship. Road one, dense and quantization-first, was the road that survived.
Method four
Weight streaming, and where the submission boundary sits
If a model's file is too big to hold entirely in RAM, one option is to page the weights in from disk as each layer needs them. We built a residency-window streaming engine for llama.cpp to test exactly that on the 2.2 GB BitCPM model, still mid-flight on road two at the time.
Development result
Throughput and RSS by resident weight budget
Stream all: 10.47 tokens per second at 279 MiB peak RSS. Pin 1,000 MB: 13.04 at 1,136 MiB. Pin 1,300 MB: 14.04 at 1,408 MiB. Pin 1,500 MB: 15.35 at 1,636 MiB. Fully resident: 18.70 at 2,129 MiB.
It worked, as engineering: full streaming cut peak RSS to 279 MiB, a huge win on paper. It also missed the 15 tok/s threshold, generating only 10.47 tok/s. Decoding one token at batch size 1 touches nearly every weight in the model, so throughput reduces to one ratio, bandwidth over model size:Hitting 20 tok/s would take about 44 GB/s of effective bandwidth. Even the fully resident point in the chart above, the fastest this model gets on this host, only reaches 41 GB/s and 18.70 tok/s. On this hardware, 20 tok/s is out of reach no matter how much RAM the residency window is given; only faster memory would clear it.
Working that ratio backward from each point in the chart above shows the same climb in different units: Stream all backs out to about 23 GB/s of effective bandwidth, Pin 1,500 MB to about 34 GB/s, Resident to about 41 GB/s. The 33 GB/s the target needs falls right at the 1,500 MB pin, where the measured rate, 15.35 tok/s, is the only point in the sweep that clears the line.
The real problem was upstream of the bandwidth math: streaming needs a custom binary, and the competition evaluates the submitted GGUF through the organiser's own unmodified llama.cpp. No engine change we make can travel with the file. That single fact drew a line we kept running into for the rest of the project:
The custom streaming engine, tensor-repacking setting, thread and KV-cache configuration, and mmap policy require a modified runtime and are not submitted. The GGUF contains the tensor quantization, tied output head, chat template, metadata, and model tensors.
Method five
Tuning the runtime before trusting any comparison
None of the model-versus-model numbers above mean anything if the runtime underneath them is inconsistent. Before comparing candidates at all, we fixed threads, KV cache, and checkpoints, and along the way tested two forms of speculative decoding that looked promising in early sweeps and fell apart under load.
Development result
Selected runtime configurations
Docker baseline: 5.3 tokens per second and 4.77 GB, still rising. Resource caps: 6.72 tokens per second and 4.44 GB. Native default: 29.78 tokens per second and 3,519 MiB physical footprint. Six threads with unified KV: 31.09 tokens per second and 3,137 MiB. Draft speculation: 24.72 tokens per second; host memory was not reliably measured.
Rejected
Draft-model speculation
Acceptance reached 98.4%, but generation still fell from 30.84 to 24.72 tok/s.
Rejected
N-gram speculation
Acceptance ran only 12–22%, so lookup and verification overhead outweighed the gain.
Adopted
Six-thread cap
Raising the thread count to ten cut decode to about 4.4 tok/s. We had no temperature reading to explain why.
Adopted
Unified KV, two checkpoints
This configuration cut retained state and sped up prefill without slowing decode.Unified KV shares capacity across active slots instead of reserving a fixed block per slot.
Six threads and unified KV got us to 31.09 tok/s at 3,137 MiB, 83% of the estimated weight-bandwidth ceiling, on the development host. None of this configuration travels with the submitted GGUF; it's the fifth method from our list, the runtime work that stays ours to keep but never enters the score. It bought us trustworthy numbers, and that finally let us run the first real model-versus-model scoreboard.
Six smaller questions
What we tried and didn't need
Six more levers, each tested once we had a trustworthy runtime to test them on. None of them changed the file we submit, but each one closed a question we'd otherwise still be asking.
GGUF-contained
Mixed tensor or layer quantization
Uniform quantization treats every tensor the same. Mixed precision keeps the parts most sensitive to rounding — embeddings, the output head, the final blocks — at a higher bit width while compressing the rest, and only pays off if the recovered accuracy is worth the extra bytes and the slower kernel those tensors fall back to. We tried it twice, on two different models: on the Qwen3 1.7B ladder, a Q3_K_M body with a Q6_K head against uniform Q3_K_M, and an IQ4_XS variant with the same higher-precision head against plain IQ4_XS; later, on Math-Expert, a Q4_0 body with a Q6_K or Q8_0 tied embedding, and separately a Q4_0 body with Q5_0 in the last four blocks.
None of it held up. The Qwen3 head swap fell to 66% ARC-Easy, worse than uniform Q3_K_M, and the IQ4_XS variant came out both slower and larger than plain IQ4_XS. Math-Expert's mixed layouts recovered no meaningful accuracy either. We rejected every mixed layout we tested — uniform quantization, chosen for the kernel it runs on rather than its nominal precision, kept winning.
GGUF-contained if supported
Structured and unstructured pruning
Structured pruning removes whole layers or low-rank factors a dense runtime can actually skip; unstructured pruning zeros individual weights, and only helps if the file format and kernel give credit for the zeros. We tested both: a Qwen control removed one layer and repeated throughput and ARC-Easy, while the BitCPM branch — covered in full in the vocabulary-pruning chapter above — measured singular-value reconstruction error on its large matrices and checked whether dense GGUF storage could benefit from sparse zeros.
Removing one Qwen layer gained about 3.7% decode speed but cost two ARC-Easy points, worth a full accuracy-score point under the scoring formula — a bad trade for a few percent of throughput. BitCPM's ternary matrices stayed full-rank, with rank-2048 factorisation error near 0.80 before quantization, and dense GGUF storage gives unstructured zeros no credit at all. We rejected every pruning branch we tested; the one pruning result we kept was vocabulary pruning, which prunes tokens rather than weights.
GGUF-contained
Smaller architectures
A smaller model moves fewer weight bytes per decoded token and usually costs less RSS — it only wins if the capability it gives up is smaller than what performance and efficiency pay back. That question isn't a side branch here; it's most of the story on this page. The four-task scale study in the second chapter, the specialist search, and the second widening further down are all, at bottom, this same question asked of a different candidate set, and all three point the same way: the 0.6B–1.7B region is where both roads, and every finalist this story ends with, actually live. See those sections for the full results.
GGUF-contained after training
Distillation and fine-tuning
Distillation transfers a larger teacher's behaviour into a smaller model; fine-tuning adapts a base model to a task — here, mathematical reasoning. Either can raise accuracy without growing the file, and Math-Expert, the specialist search's raw leader, is exactly this kind of result: a Qwen3-0.6B fine-tune on OpenMathReasoning-mini. We screened several more finished public derivatives in that same search (its provenance section records why each one that didn't make the cut fell short), but ran no new training of our own.
We kept Math-Expert as a finished, tested fine-tune, and deferred new training rather than submit something we couldn't fully validate: a GPU, a full dataset audit, and a reproducible conversion path were all missing pieces we couldn't fill in the time we had.
Stored layout plus engine-only policy
Tensor layout, alignment, and runtime repacking
GGUF fixes how tensors are stored on disk, but the engine can still repack them into a faster in-memory layout at load time. Repacking can speed up a kernel, but it costs the RSS budget to do it. A 4B development control compared the default runtime with repacking disabled; we also reviewed a custom alignment or packing scheme on top of that, and decided against building one.
Disabling repacking cut the tested 4B footprint from about 3,236 to 602 MiB, a real product win, without a clear speed loss. But an unsupported layout that fails to load doesn't get a partial score, it gets zero, so we kept the stock GGUF layout for compatibility rather than risk it. No-repack stays a product-only setting — it's method ten on our list, the boundary this whole story keeps running into.
Engine-only
Context size and KV-cache implications
Context length sets how many token states the KV cache holds. A shorter or quantized cache can save memory, but it doesn't touch the weight traffic a batch-one decode has to move every token regardless of context. We'd already replaced runtime defaults with explicit context and KV limits back in the runtime-tuning chapter, which stopped memory growth and made later comparisons trustworthy — and the participant profiler fixes its own workload anyway, at a 512-token prompt and 128 generated tokens.
No GGUF-level context metadata we could set would change what gets scored, so we rejected context metadata as a profiler-facing lever. We keep the bounded context and KV settings we already adopted, but only as product configuration.
Checkpoint
The first scoreboard: Qwen3 1.7B wins, narrowly
With quantization, deduplication, and the runtime settled, we ran six candidates through the actual participant profiler. Qwen3 1.7B Q4_0, pure and tied, won at 72.4653, ahead of Qwen3.5 0.8B by 0.92 points and comfortably ahead of both the BitCPM ternary bet and a larger Qwen3.5 4B candidate.
Official profiler result
Score by component
Official profiler result
Speed against memory
Scalar measurements
| Model | tok/s | First token | Peak RSS | ARC-Easy | Total |
|---|---|---|---|---|---|
| Math-Expert 0.6B Q4_K_MScalar leader · 396.7 MB | 12.72 | 23.61 s | 540 MiB | 68% 54.2–79.2 | 77.9324 |
| Qwen3.5 0.8B Q4_0Risk-adjusted recommendation · 507.2 MB | 12.63 | 16.62 s | 670 MiB | 64% 50.1–75.9 | 75.3895 |
| Qwen3 1.7B Q4_0Previous scalar leader · 974.2 MB | 9.79 | 35.37 s | 1,116 MiB | 72% 58.3–82.5 | 72.4653 |
| Qwen3.5 0.8B Q4_K_M | 9.74 | 28.26 s | 695 MiB | 68% 54.2–79.2 | 71.54 |
| BitCPM4 8B TQ2_0 | 0.81 | 584.15 s | 2,307 MiB | 88% 76.2–94.4 | 59.18 |
| Qwen3.5 4B IQ4_XS | 1.13 | 395.19 s | 2,627 MiB | 76% 62.6–85.7 | 52.93 |
The final Qwen model has the shortest measured first-token latency, 16.62 seconds against Math-Expert's 23.61. Neither number enters the score.
The table above already shows Math-Expert and the final Qwen3.5 0.8B ahead of the 1.7B model that won this first round. That's the ledger as it stands today: two later rounds moved the winner twice more. At the time of this first scoreboard, though, 1.7B genuinely led, and we treated it as settled long enough to ask a harder question: would the public webpage's cohort-relative scoring formula, instead of the fixed 15, change anything?
We checked, and the full breakdown, at every cohort floor we tested, is in the operational appendix under Website-relative sensitivity. It's a separate check, never blended with the profiler result above.
Method six
Finding the real CPU, and watching the ranking flip
Everything so far assumed the audit binary runs scalar only, because that's how the executable profiler we could get our hands on was built. But the competition's own public page describes hardware that plausibly supports the vector configuration defined above. We couldn't be sure which one the final judging actually runs, so we built both and measured the same seven models on each.
AVXON
AVX2ON
FMAON
F16CON
NATIVEOFF
AVX-512OFF
Matched models and scoring function
Total score by CPU method
Loading paired scalar and vector score evidence.
Turning the vector configuration on reordered the field as well as sped it up. Q4_K_M, penalised by the generic-C fallback under quantization's first lesson, jumped ahead of pure Q4_0 once it had a real vector kernel to run on: 80.4484 against 80.2818, a lead earned entirely by 59.7 MiB less repacking RSS, not by anything about the model, and not by any change to its measured accuracy. Even BitCPM, road two's abandoned bet, sped up 9.235× and became operationally viable, though its total still couldn't catch the Qwen variants.
Scalar and vector measurements
| Model | Scalar → vector pp512 | Scalar → vector tg128 | Decode gain | Est. profiler RSS | ARC-Easy proxy | Total |
|---|---|---|---|---|---|---|
| Math-Expert 0.6B Q4_K_MExpanded-search leader | 21.7268 → 153.9351 | 12.6339 → 39.2320 | 3.105× | 759.7 MiB+202.9 MiB | 68% | 81.8803 |
| Qwen3.5 0.8B Q4_0Risk-adjusted recommendation; tensor-identical vector source | 31.0781 → 98.0094 | 12.6955 → 27.1509 | 2.139× | 928.1 MiB+240.8 MiB | 64% | 79.4104 |
| Qwen3 1.7B Q4_0Previous scalar choice | 14.7166 → 47.0716 | 9.9869 → 16.8927 | 1.691× | 2,049.4 MiB+916.3 MiB | 72% | 80.2818 |
| Qwen3 1.7B Q4_K_MQuantization-ladder leader | 7.1859 → 55.4554 | 5.2954 → 15.6714 | 2.959× | 1,989.7 MiB+806.2 MiB | 72% | 80.4484 |
| Qwen3 1.7B Q5_K_M | 6.5613 → 24.3231 | 4.7839 → 12.7191 | 2.659× | 1,364.6 MiB+0.1 MiB | 76% | 79.6307 |
| Qwen3 1.7B IQ4_XS | 3.2063 → 23.9364 | 2.4961 → 14.0644 | 5.635× | 1,082.3 MiB+0.5 MiB | 70% | 80.1089 |
| BitCPM4 8B TQ2_0 | 0.8762 → 13.6569 | 0.8108 → 7.4876 | 9.235× | 2,316.4 MiB+0.1 MiB | 88% | 72.5121 |
Method seven
Widening the search finds a specialist we hadn't considered
Once the Qwen roster stopped producing gains, we widened the net: nine more candidates, and eight quantization layouts of one promising fine-tune, Math-Expert, a Qwen3-0.6B model tuned on mathematical reasoning data. It won the raw scoreboard. It didn't win cleanly.
Profiler proxy
Math-Expert 0.6B Q4_K_M
- Direct scalar total
- 77.9324
- Vector total, n=50
- 81.8803
- Vector generation
- 39.23 tok/s
- Est. vector RSS
- 759.7 MiB
Use this file only if you treat the executable profiler's 50 ARC-Easy items as the selection objective.
Risk-adjusted submission
Muta-Tutor-Qwen3.5-0.8B-Q4_0-final.gguf
- Direct scalar total
- 75.3895
- Vector total, n=500
- 76.8104
- Vector generation
- 27.15 tok/s
- Est. vector RSS
- 928.1 MiB
This file wins the matched result outside the 50-item profiler slice and reaches its first token sooner. Its template supplies the tutor policy and disables hidden reasoning whenever the caller sends no setting of its own.
Latest finalists, fixed-15 score
Direct scalar and vector totals
Loading the latest finalist vector comparison.
Latest scalar and vector finalist results
ARC-Easy-50 supplies the raw fixed-15 comparison; the ARC-Easy-500 column is diagnostic. The final Qwen vector row transfers the source measurement, and only after we verified tensor identity.
Sample sensitivity
Profiler total and larger-sample diagnostic
Loading the finalist score comparison.
Quantization search
Speed and ARC-Easy trade-off
Loading the Math-Expert quantization sweep.
We ran Math-Expert through the same quantization-method question we'd already answered once, on a completely different model, and got the same answer: the format with a real kernel wins, and no amount of extra precision in the embedding or final layers buys back what a slow kernel costs. That's the second confirmation we promised back in the quantization section.
Finalist accuracy checks
Confidence intervals are Wilson 95% intervals. These task results are diagnostics; they are not the judging-panel tutoring score.
Models screened and rejection reasons
Everything else we tried in this round
| Method | Test or constraint | Decision |
|---|---|---|
| Uniform quantization | Q4_0, Q5_0, Q4_K_S, Q4_K_M, and IQ4_XS from the same F16 source. | Keep Q4_K_M. Q4_0 is fast but loses too much accuracy; Q5_0 is slow on the scalar profiler kernel. |
| Mixed tensor precision | Q4_0 body with Q6_K or Q8_0 tied embedding; Q5_0 in the final four blocks. | Reject. None recovered enough ARC-Easy accuracy. |
| Architecture and specialist search | Qwen3/3.5, Gemma 3, Noema, VibeThinker, OpenMath-Nemotron, NuminaMath, and a reasoning distill. | Keep Qwen3.5-0.8B and Math-Expert as finalists. Other files lost on throughput, accuracy, licensing, or provenance. |
| Pruning and vocabulary reduction | Prior layer-pruning and BitCPM vocabulary results were retained; Qwen BPE vocabulary pruning cannot be reproduced safely with the current tool. | No new submission candidate. |
| Distillation and fine-tuning | Finished public fine-tunes were tested. A new training run was not promoted without a GPU, a complete dataset audit, and a reproducible conversion path. | Defer new training; do not submit a partially validated derivative. |
| MTP and speculative decoding | llama-bench in the profiler does not invoke the runtime flags required by MTP or draft decoding. | Exclude from a GGUF-only profiler claim. |
| Context and embedded template | The profiler fixes p512/tg128. A live four-prompt test showed that unrestricted hidden reasoning could consume the full response allowance. The final GGUF embeds the tutor policy and forces non-thinking ChatML. | Adopt for interactive reliability; assign no profiler throughput gain. |
Research and provenance
We built the search on public model cards, documented licences, reproducible conversion procedures, and retained raw measurements. Math-Expert is a Qwen3-0.6B fine-tune on OpenMathReasoning-mini; its public card gives no complete training recipe or benchmark table. The Qwen3.5 model derives from the official Qwen checkpoint.
Primary references: Math-Expert model card, GGUF repository, OpenMathReasoning dataset, OpenMathReasoning paper, low-bit reasoning study, and llama.cpp speculative-decoding documentation.
Method seven, continued
A second widening moves the vector leader again
Math-Expert's win over the six-model ledger settled nothing about architecture, only about one specialist fine-tune. We hadn't tried Qwen's own general-purpose 1.5B and 3B line at all, or comparable dense models from outside the Qwen tree. So we paired eight GGUFs — Qwen2.5 at 1.5B and 3B, Qwen2 at 1.5B, Llama 3.2 at 1B and 3B, Gemma 2 at 2B, Phi-4 Mini, and Orca Mini at 3B — and ran every one of them through the same 512-token prompt and 128-token decode, under both the scalar and vector configurations, with each model's accuracy measured exactly once and reused for both totals.
Eight matched GGUFs
Total score by CPU configuration
Loading paired scalar and vector score evidence for the eight-model screen.
Qwen2.5 1.5B took first under vector, 82.8697, with Qwen2 1.5B close behind at 82.2176 — the two smallest, most conventionally quantized entries in the field. Llama 3.2 1B placed third at 78.0999 on the least memory of any candidate, 1,397.8 MiB. Gemma 2 2B, quantized to Q8_0 instead of Q4_K_M, repeated quantization's oldest lesson from a new angle: its decode gain under vector was only 1.826×, against roughly 3× for the Q4_K_M models, because Q8_0 doesn't reach the same SIMD path. 74% accuracy wasn't enough to make up the difference, and it finished at 61.9274. Orca Mini 3B saw the largest decode gain of any model in this story, 6.583×, but started so slow under scalar and scored so low on accuracy, 58%, that vector execution only lifted it to 52.4683.
Eight-model scalar and vector measurements
| Model | Scalar → vector pp512 | Scalar → vector tg128 | Decode gain | Vector RSS | ARC-Easy, n=50 | Total |
|---|---|---|---|---|---|---|
| 1 · Qwen2.5 1.5B Q4_K_M | 7.6824 → 59.9786 | 5.6919 → 17.2265 | 3.026× | 1,838.7 MiB | 76% | 65.9176 → 82.8697 |
| 2 · Qwen2 1.5B Q4_K_M | 7.7000 → 59.1846 | 5.6976 → 16.8156 | 2.951× | 1,714.0 MiB | 74% | 65.2782 → 82.2176 |
| 3 · Llama 3.2 1B Q4_K_M | 10.6776 → 82.0402 | 7.0795 → 21.7591 | 3.074× | 1,397.8 MiB | 64% | 63.5027 → 78.0999 |
| 4 · Qwen2.5 3B Q4_K_M | 3.6584 → 28.4679 | 2.9188 → 8.8583 | 3.035× | 3,470.9 MiB | 78% | 58.6863 → 67.0322 |
| 5 · Llama 3.2 3B Q4_K_M | 3.6258 → 28.1515 | 2.8118 → 8.6254 | 3.068× | 3,454.6 MiB | 72% | 55.6109 → 63.6118 |
| 6 · Gemma 2 2B Q8_0 | 5.5196 → 28.7048 | 3.5435 → 6.4692 | 1.826× | 2,871.1 MiB | 74% | 56.0770 → 61.9274 |
| 7 · Phi-4 Mini 3.8B Q4_K_M | 3.1238 → 19.3956 | 2.3084 → 6.6995 | 2.902× | 3,827.2 MiB | 78% | 56.3229 → 61.7203 |
| 8 · Orca Mini 3B Q4_K_M | 0.8970 → 13.0483 | 0.8482 → 5.5836 | 6.583× | 2,759.3 MiB | 58% | 42.9981 → 52.4683 |
Narrowing. Promote Qwen2.5 1.5B as the new vector candidate; eliminate the other seven from that shortlist. None of this transfers to the scalar profiler, where the same Q4_K_M kernels that helped under vector are the slow generic-C path quantization already taught us to avoid.
76% on 50 ARC-Easy items is a small-sample estimate, and Qwen2.5 1.5B was about to become our leading candidate on the strength of it, so we ran the same benchmark again at ten times the sample size. The result: 71.8%, with a 95% interval of 67.70–75.57%. The scoring function takes accuracy as one input alongside throughput and RSS, which we already had measured and held fixed, so both totals move by exactly the same 2.10 points the accuracy alone accounts for: the vector total falls to 80.7697, the scalar total to 63.8176. Nothing about the CPU kernel changed; only the sample size did.
| Qwen2.5 1.5B Q4_K_M | n=50 | n=500 |
|---|---|---|
| ARC-Easy accuracy | 76.0% | 71.8% (67.70–75.57% CI) |
| Scalar total | 65.9176 | 63.8176 |
| Vector total | 82.8697 | 80.7697 |
That gave us reason to hold Math-Expert and the final Qwen3.5 file to the same larger sample. Both already had a matched 500-item ARC-Easy run on file from the specialist search: Qwen3.5 at 58.8%, Math-Expert at 54.6%, well below either one's 50-item read. Lined up against Qwen2.5's own 500-item score, the picture splits cleanly by CPU configuration:
| Finalist, matched to n=500 | ARC-Easy, n=500 | Scalar total | Vector total |
|---|---|---|---|
| Qwen2.5 1.5B Q4_K_MVector leader | 71.8% 67.70–75.57 | 63.8176 | 80.7697 |
| Qwen3.5 0.8B Q4_0 finalScalar leader | 58.8% 54.43–63.03 | 72.7895 | 76.8104 |
| Math-Expert 0.6B Q4_K_M | 54.6% 50.22–58.91 | 71.2324 | 75.1803 |
Under vector, Qwen2.5 1.5B leads all three finalists. Under scalar, the ranking reverses entirely: Qwen3.5 leads, Math-Expert follows, and Qwen2.5 falls to third, because the profiler kernel that runs Qwen2.5's Q4_K_M tensors is the same slow generic-C path that cost every k-quant in this story. Neither raw finalist we found first turns out to be the best answer in every regime; the honest one depends on which binary the final audit actually runs.
The combined figure retains the original 15-model search and adds the two tuned finalists, so the latest result remains visible beside the candidates that produced it:
Model search and tuned finalists
Every paired model, scalar versus vector
Loading the combined comparison across both explorations.
Narrowing. Qwen2.5 1.5B becomes the vector-configuration recommendation, provisionally, pending the same physical-target run and embedded-behaviour check Qwen3.5 has already been through. Qwen3.5 0.8B remains the scalar-configuration recommendation, ahead of Math-Expert once both are read at the same sample size. The submission choice depends on which CPU configuration the supplied profiler actually runs, not on either finalist's small-sample score alone.
Method eight
Teaching the file to behave, since nothing else is in the room
Every method so far shrinks bytes or speeds up decode. This one changes what the model says: judges chat with the bare GGUF directly, no system prompt and no application layer standing between them and the model, so the tutoring persona has to travel inside the file or it doesn't exist during evaluation.
The Qwen3.5 0.8B Q4_0 source is 507.2 MB. The build adds the tutor policy, English language tag, non-thinking ChatML template, and sampling defaults without changing model tensors. Verification confirms that all 320 model tensors are identical.
We replace the chat template with a clean ChatML template that injects the Muta tutoring persona as the system turn whenever the client sends none, and always opens the assistant turn with an empty <think></think> block, forcing a direct answer instead of a long reasoning trace. That last part isn't cosmetic. A live four-prompt acceptance test with unrestricted thinking left both finalists spending their entire 256-token allowance on hidden reasoning and returning no usable answer at all, a minute of silence per prompt at audit-box speed. Disabling thinking fixed that. Sampling defaults (temperature 0.4, top-p 0.9, min-p 0.05, repeat penalty 1.05) are read straight from the file by llama-server, and the persona runs about 130 tokens, so a judge's first turn pays little extra prefill for it.
We verified this on llama.cpp's Jinja engine and llama-cpp-python's separate jinja2 path. The previous Qwen3.5 package injected the persona and stopped correctly, but the hardest proof prompt still failed. The fine-tuned finalists below have not yet repeated this battery; measured accuracy alone does not clear that requirement.
Method nine
Fine-tuning the two finalists
We trained 15 candidates across Qwen3.5 0.8B and Qwen2.5 1.5B, varying LoRA rank, BF16 LoRA versus QLoRA, data mixture, learning rate, and training length. Every survivor was exported to its deployment quantization and evaluated as a GGUF; training loss did not select a winner.
The first eight runs did not improve ARC-Easy-500. Most training examples were long solutions or tutor dialogue, while the profiler scores short answer continuations. The original data builder also omitted the leading space between Answer: and the candidate continuation, changing the BPE tokens used during training. We corrected both problems, filtered against held-out questions, and used only source training splits. The final Qwen2.5 mixture excludes OpenBookQA because its published licence is unclear.
Initial sweep
Eight runs · no promotion
Balanced and reasoning-heavy mixtures improved validation loss but did not improve the exact profiler task.
Metric-aligned sweep
Seven runs · two promotions
Raw multiple-choice continuations, corrected token boundaries, verified answers, and leakage checks improved both finalists.
Matched 500-item evaluation
Accuracy before and after tuning
Loading fine-tuning accuracy results.
Fine-tuned GGUFs
Total score by CPU configuration
Loading fine-tuned finalist scores.
Matched control, performance, and memory results
Qwen3.5 gained 15.0 accuracy points and approximately 7.5 total points in either CPU configuration. Qwen2.5 gained 3.4 accuracy points and approximately 1.7 total points. Throughput and memory were effectively unchanged within each matched pair, so these score gains come from accuracy rather than a deployment trade-off.
Secondary held-out checks
Current decision
Fine-tuning raises both scores; the runtime split remains
The matched 500-item comparison still selects one model per CPU configuration. Fine-tuning raises both totals without materially changing speed or memory.
Scalar configuration, n=500
Fine-tuned Qwen3.5 0.8B · 80.3664
70.2% ARC-Easy, 13.60 tok/s, and an estimated 691 MiB profiler RSS.
Vector configuration, n=500
Fine-tuned Qwen2.5 1.5B · 84.1387
77.8% ARC-Easy, 17.44 tok/s, and an estimated 1,706 MiB profiler RSS.
Quantization decides the available CPU kernel; fine-tuning only changes accuracy on top of it. So the final choice still comes down to which CPU configuration the evaluator actually runs.
The receipts
Every experiment, adopted, rejected, and deferred
The measurements above produced the decisions below. Adopted changes appear first; use the filters to see rejected, neutral, and deferred tests.
Challenge requirements
Requirement status
Model selection also has to satisfy the published challenge requirements. Each entry pairs one requirement with what we have actually implemented.
Challenge requirements follow the Africa Deep Tech Challenge 2026 page.
Operational appendix
Profiler results and comparison tables
The underlying measurements, live profiling controls, stored results, and mixed-machine archive all stay available below.
Published snapshot. This read-only copy was rendered from the repository’s stored evidence. Profiling, promotion, and record deletion need the local dashboard server (./dashboard/start.sh).
Official profiler resultDirect GGUF campaignLoading campaign evidence…
Profiler-parity estimateCandidate screening with the scalar audit kernelLoading reconstructed profiler evidence…
Controlled vector proxyPaired scalar and vector measurements, seven-model ladderLoading the seven-model instruction-set comparison…
Controlled vector proxyPaired scalar and vector measurements, eight-model architecture screenLoading the eight-model architecture screen…
Website-relative sensitivityVector measurements with the cohort-relative formulaLoading sensitivity data…
Profiling…
Mixed-machine archiveHistorical profiler archiveDifferent machines and engine configurations; excluded from the campaign ranking.
Historical display only: S_total = 0.50·S_acc + 0.30·S_perf + 0.20·S_eff − 10·thermal · legacy S_perf = min(TPS / TPS_max, 1) × 100 · TPS_max = fastest stored run · S_eff = (7 GB − peak RAM) / 7 GB × 100 · crash or OOM ⇒ disqualified (0).