Experiment report Loading profiler…

The story of a competition entry

Teaching a model to teach

I picked Math & Scientific Reasoning because pragmatic education is something African classrooms need and rarely get enough of: too many students for one teacher, and too little time to slow down on the question any single student actually got wrong.

The competition evaluates an offline model on a $150–$500 CPU-only laptop. The model must therefore balance teaching accuracy, response speed, and memory use. This report follows the experiments that narrowed those three variables to two runtime-dependent finalists.

Where the story stands, for now

Fine-tuning improved both finalists

Qwen3.5 0.8B remains the scalar leader after a 15-point ARC-Easy gain. Qwen2.5 1.5B remains the vector leader after a smaller 3.4-point gain. These are the best measured GGUFs, not yet the packaged submission: both still need the tutor metadata and live-prompt checks, and the final choice still depends on the audit CPU configuration.

Scalar leader, n=500Qwen3.5 0.8B Q4_0

Scalar total
80.3664
Vector total
82.3962
Scalar generation
13.60 tok/s
Est. scalar RSS
691 MiB
ARC-Easy, n=500
70.2%

Vector leader, n=500Qwen2.5 1.5B Q4_K_M

Vector total
84.1387
Scalar total
67.0475
Vector generation
17.44 tok/s
Est. vector RSS
1,706 MiB
ARC-Easy, n=500
77.8%

Selection condition: use the tuned Qwen3.5 candidate for the supplied scalar configuration; use tuned Qwen2.5 only if the audit is confirmed to use the vector configuration. Before submission, apply the same embedded tutor policy to the chosen GGUF and repeat the live-prompt and physical-target checks.

Scalar candidate size
513.0 MB
Quantization
Q4_0
Target
8 GB, CPU-only
Selection basis
Matched n=500 accuracy, speed, and RSS
01

The rules we're playing by

A $150–$500 laptop is the whole target

The Africa Deep Tech Challenge wants language models running entirely on-device, on the standard laptops already sitting in African cities, with no cloud behind them. It lists seven tracks: Math & Scientific Reasoning, Healthcare, Agriculture, Creative Writing, Coding Assistants, Corporate/Enterprise, and Autonomous Agents. We took the first one, because a tutor that explains a wrong answer instead of just marking it is closer to what a scarce teacher actually does.

The hardware is exact: an Intel Core i5, 10th to 12th generation, integrated graphics only, Ubuntu 22.04, and a hard 7 GB RAM ceiling. Go over it and the run is disqualified, full stop. The rules are just as narrow on software: llama.cpp only, weights in GGUF. Any open base model is fair game, quantized or fine-tuned however we like, but nothing closed and nothing needing a different runtime.We disable AVX-512 in our own build and check the disassembly for it, because several eligible consumer CPUs don't support it: an illegal instruction there means an automatic zero.

Stotal= 0.50·Sacc+ 0.30·Sperf+ 0.20·Seff− Pthermal
Fig. 1 Fixed-15 profiler score. Below 15 tok/s, an increase of 1 tok/s adds 2.00 total points. One accuracy point adds 0.50. A reduction of 1 GB in peak RSS adds 2.86.

Half the score is accuracy, so a fast model that can't teach still loses. But performance caps at 15 tok/s and efficiency zeroes out at 7 GB, so past those points speed and memory stop paying at all.S_perf = min(TPS/15, 1) × 100 and S_eff = max(0, (7 GB − RSS)/7 GB) × 100. The public challenge page instead defines performance relative to the fastest submission, TPS_max, rather than a fixed 15. We follow the fixed formula, since it's what the executable profiler actually runs, and keep the cohort-relative reading as a separate sensitivity check throughout this page. That formula produces a decision rule we checked every optimisation against for the rest of the project:

Two CPU configurations run through this story, named once so we can stop repeating the acronyms. The scalar configuration is the supplied profiler build, with the wider vector extensions disabled for the affected kernels. The vector configuration is a portable SIMD build with AVX2, FMA, and F16C enabled, on the same two physical cores. From here on, this page just says scalar and vector.

Why we keep five separate evidence categories on this page

Our numbers come from different systems, CPU features, and performance formulas, so we never blend them into one ranking: official profiler results (the participant executable's own runs), profiler-parity estimates (a scalar-only reconstruction with a 45 MiB profiler-root allowance we estimated, not measured), controlled vector proxies (a paired scalar/vector benchmark on the same host), website sensitivity (the cohort-relative formula as a separate check, never averaged in), and plain development results (Mac, Docker, and custom-engine tests that guided decisions but never ranked a submission). A colour and a badge on every figure below show which lane it belongs to.

02

Before we picked a model

Bigger teaches better. Bigger also costs more.

Before touching quantization, we ran a plain accuracy-by-scale study across the Qwen3.5 family: 0.8B, 2B, and 4B, on four reasoning tasks. The result was the least surprising thing in this whole project, and also the thing every later decision had to argue against.

Separate accuracy evaluations

Candidate measurements

Reasoning studyAudit campaign

The separate four-task reasoning study recorded Qwen3.5 4B at 2.55 GiB and 73.3, 2B at 1.19 GiB and 64.8, and 0.8B at 0.50 GiB and 51.3. The audit-candidate panel uses separate ARC-Easy proxies. Fine-tuned Qwen3.5 0.8B is 0.48 GiB at 70.2% and leads the scalar total. Fine-tuned Qwen2.5 1.5B is 0.92 GiB at 77.8% and leads the vector total. Earlier candidates remain as comparison points.

Fig. 2 The reasoning-study points use a four-task development mean; audit candidates use ARC-Easy proxies. The two fine-tuned finalists are marked separately from the historical candidates.

4B beat 2B beat 0.8B, cleanly, on every task. Under a formula that pays 0.50 total-score points per accuracy point and only 2.86 per gigabyte saved, that curve makes a strong case for going as big as the RAM ceiling allows. But RSS on Linux tracks file bytes almost one-to-one, so "as big as the ceiling allows" is a number in the hundreds of megabytes, not gigabytes, once quantization is priced in. That tension split our search down two roads at once, plus one number neither road was allowed to cross:

Road one

Go small and dense

Stay in the Qwen3/3.5 family, quantize as hard as the audit kernel would tolerate, and win back accuracy through parameter dedication and behaviour baked into the file rather than raw scale.

Road two

Go big and ternary

Bet that an 8B-parameter model at roughly 2 bits per weight, BitCPM4 in TQ2_0, could out-accuracy anything dense enough to fit, if we pruned and compressed it hard enough to land in the same weight class as the dense candidates.

Every 100 MB above the smallest working file was already worth 1.4 points of efficiency score we'd rather spend on accuracy, so neither road got a free pass on size: road two's ternary bet only earned its keep by landing at about 2.1 GB, not the 8 GB the parameter count alone would suggest. We ran both roads in parallel for weeks. This page tells road one's story in full, because it's the one that shipped; road two still shows up wherever its lessons apply, and its ending is in the vocabulary-pruning section below.

03

The toolbox

Ten optimization methods

We tested ten ways to change accuracy, file size, RAM, or speed. Each method either narrowed the candidate set or was retained as a measured rejection.

1

Quantization

Which quantization type the audit binary can actually run fast, not just how many bits per weight.

2

Parameter deduplication

Dropping the duplicated output head when embedding and head can share one matrix.

3

Vocabulary pruning

Removing tokens the tutor will never emit, and checking token by token that nothing else changed.

4

Weight streaming

Paging weights from disk instead of keeping the whole file resident, and finding the boundary the competition draws around it.

5

Runtime tuning

Threads, KV cache, checkpoints, and two kinds of speculative decoding, before any model comparison could be trusted.

6

Instruction-set targeting

Discovering that the audit binary's actual CPU features change which quantization type wins, and testing both readings.

7

Widening the model search

Testing nine more candidates and a specialist fine-tune once the first roster stopped improving.

8

Baking behaviour into the file

Persona, chat template, and sampling defaults, since judges talk to the bare GGUF with no app layer in front of it.

9

Fine-tuning

Updating the two finalists on verified math and science questions matched to the profiler's continuation format.

10

Runtime configuration we couldn't submit

Repacking, mmap policy, thread count: measured changes that live in Muta's runtime, not in the submitted GGUF.

Quantization, deduplication, embedded behaviour, and fine-tuning can change the submitted GGUF. Streaming and runtime tuning cannot. The model search identifies which architecture receives those changes; the scalar/vector comparison determines how each candidate is scored.

04

Method one

Quantization: the file format is the kernel choice

Quantization means storing each weight in fewer bits than the 16 or 32 it was trained in, trading a little precision for a much smaller file. GGUF supports a whole ladder of these formats, from plain 4-bit rounding (Q4_0) through vector-quantized "k-quants" (Q4_K_M, Q5_K_M) to sub-4-bit "i-quants" and ternary formats. We assumed, like most guides do, that a newer, smarter format like Q4_K_M would be the better choice. On this evaluator, it wasn't.

Scalar profiler configuration

Kernel dispatch by tensor type

Scalar audit build
Tensor quantisation type determines the available CPU kernel path Scalar configuration · vector extensions disabled Tensor type in the GGUF Available CPU kernel Measured decode Q4_0 current Qwen final Hand-written SSSE3 kernel vectorised, 16 bytes per instruction 12.63 tok/s current Qwen final Q4_K_M, Q5_K_M IQ4_XS TQ2_0, TQ1_0 k-quants, i-quants, ternary Generic C fallback scalar, one weight at a time 0.81–12.72 tok/s current six-model ledger Source analysis maps Q4_0 to SSSE3 and the other listed types to generic C; campaign rows show model-level rates.

On the scalar participant build, Q4_0 tensors use a hand-written SSSE3 kernel and the current Qwen final decodes at 12.63 tokens per second. K-quants, i-quants, and ternary types use scalar generic C implementations; their measured rates span 0.81 to 12.72 tokens per second across the current six-model ledger because model scale also differs. The vector build enables SIMD kernels for these tensor types and changes the quant ordering, without changing any model's measured accuracy.

Fig. 3 We show the current scalar-ledger rates beside the kernel paths. Model size also affects the measured range. The vector build enables SIMD implementations for the listed tensor types.

We read the audit binary's own kernel source rather than trust a guide written for a different CPU, and found only Q4_0 has a hand-written SIMD path on the scalar build the executable profiler runs. Everything else, including the "smarter" k-quants, falls back to generic C. The original control, on the Qwen3 1.7B pair, measured 9.99 tok/s in Q4_0 against 5.30 tok/s in Q4_K_M, purely from that one format choice; the diagram above shows the same pattern holding on the file we eventually shipped, at 12.63 tok/s. That single fact ruled out most of the usual advice and pointed every candidate toward pure Q4_0, at least until instruction-set targeting complicated the picture again, several sections down.

05

Method two

Cutting the duplicate head

Many transformer architectures give the embedding table and the output head separate weight matrices, even though both are the same shape and often learn near-identical representations. Tying them, letting the head reuse the embedding matrix instead of storing its own copy, is a free win if the architecture supports it: no retraining, no accuracy cost, just a smaller file.

We reproduced our recommended file from its source quant using metadata changes alone, then ran an isolated A/B on the tied-versus-untied question: removing the duplicate saved about 175 MB with identical ARC-Easy accuracy in both conditions, an unambiguous win we took on every candidate that supported it.

The tutor persona, chat template, and sampling defaults go in through a similar scripted pass — though that one only touches metadata, not tensors, so it costs no file size at all. It gets its own full treatment in the behaviour-baking section further down, once the mechanisms that shrink the file are out of the way.

06

Method three, and the end of road two

Pruning the vocabulary, and letting the ternary bet go

Vocabulary pruning removes tokens a model will never realistically emit from its embedding table, shrinking the largest single matrix in a small model without touching a single weight the tutor actually uses. We built this for road two, the 8B ternary BitCPM4 model, since an English tutor has no use for most of a 73,448-token vocabulary built to cover CJK scripts.

01

Vocabulary pruning

We cut the vocabulary from 73,448 to 44,416 tokens and padded it to a multiple of 64, which shrank the file by 164 MB. English tokenisation matched across 20,464 checked tokens; perplexity fell from 10.558 to 10.473.

Kept
02

TQ1_0 body test

The body shrank by 340 MB, but generic CPU throughput fell from 3.70 to 2.88 tok/s. The lower bit width bought us nothing on the evaluated kernel's instruction cost.

Rejected
03

Head and embedding quantization

Head and embedding requantization saved at most 48 MB. The largest estimated total-score gain we found was about 0.14 points, too small to matter.

Rejected
04

Factorisation and sparsity tests

The ternary matrices turned out to be full-rank: rank-2048 factorisation error was about 0.80 before quantization. Dense GGUF storage and kernels give us no credit for unstructured zeros either, so sparsity saved nothing.

Rejected

Only the vocabulary prune was worth keeping, and it was the last real win road two got. BitCPM4-8B-TQ2_0 still recorded the highest accuracy of any candidate we ever tested, 88% ARC-Easy, well above every Qwen quantization we ran. But it decoded at just 0.81 tok/s on the scalar audit kernel, since ternary formats fall on the same generic-C path as k-quants. Even after instruction-set targeting later gave it a vector kernel and lifted that to 7.49 tok/s, its total reached only 72.5121, still short of every Qwen variant. The most accurate model we built turned out to be the one we couldn't ship. Road one, dense and quantization-first, was the road that survived.

07

Method four

Weight streaming, and where the submission boundary sits

If a model's file is too big to hold entirely in RAM, one option is to page the weights in from disk as each layer needs them. We built a residency-window streaming engine for llama.cpp to test exactly that on the 2.2 GB BitCPM model, still mid-flight on road two at the time.

Development result

Throughput and RSS by resident weight budget

Custom engine

Stream all: 10.47 tokens per second at 279 MiB peak RSS. Pin 1,000 MB: 13.04 at 1,136 MiB. Pin 1,300 MB: 14.04 at 1,408 MiB. Pin 1,500 MB: 15.35 at 1,636 MiB. Fully resident: 18.70 at 2,129 MiB.

Fig. 4 A historical custom-runtime experiment, completed before the current model campaign. Each point pins more of the 2.2 GB BitCPM model in memory; the result is not a submission score.

It worked, as engineering: full streaming cut peak RSS to 279 MiB, a huge win on paper. It also missed the 15 tok/s threshold, generating only 10.47 tok/s. Decoding one token at batch size 1 touches nearly every weight in the model, so throughput reduces to one ratio, bandwidth over model size:Hitting 20 tok/s would take about 44 GB/s of effective bandwidth. Even the fully resident point in the chart above, the fastest this model gets on this host, only reaches 41 GB/s and 18.70 tok/s. On this hardware, 20 tok/s is out of reach no matter how much RAM the residency window is given; only faster memory would clear it.

tok/s≈ BWeffM
Fig. 5 Effective bandwidth blends fast reads from whatever's resident with slow reads from whatever still has to stream off the SSD. Reaching 15 tok/s on this 2.2 GB model needs about 33 GB/s of it.

Working that ratio backward from each point in the chart above shows the same climb in different units: Stream all backs out to about 23 GB/s of effective bandwidth, Pin 1,500 MB to about 34 GB/s, Resident to about 41 GB/s. The 33 GB/s the target needs falls right at the 1,500 MB pin, where the measured rate, 15.35 tok/s, is the only point in the sweep that clears the line.

The real problem was upstream of the bandwidth math: streaming needs a custom binary, and the competition evaluates the submitted GGUF through the organiser's own unmodified llama.cpp. No engine change we make can travel with the file. That single fact drew a line we kept running into for the rest of the project:

Runtime-only changes and GGUF-contained changes Requires a modified runtime Not included in the model-only submission Residency streaming engine 2,129 → 279 MiB Disable tensor repacking 3,236 → 602 MiB Thread and KV-cache configuration development-runtime settings Lazy mmap policy engine setting Encoded in the submitted GGUF Evaluated by the organiser's runtime Q4_0 or Q4_K_M tensor layout runtime kernels determine measured cost Tied output head approximately 175 MB smaller in A/B test Chat template and metadata persona and sampling defaults Model tensors unchanged from source quant The submitted file cannot alter llama.cpp build flags, memory policy, thread count, or cache configuration.

The custom streaming engine, tensor-repacking setting, thread and KV-cache configuration, and mmap policy require a modified runtime and are not submitted. The GGUF contains the tensor quantization, tied output head, chat template, metadata, and model tensors.

Fig. 6 We classify every evaluated change by submission eligibility here. Only properties encoded in the GGUF can affect the model-only evaluation.
08

Method five

Tuning the runtime before trusting any comparison

None of the model-versus-model numbers above mean anything if the runtime underneath them is inconsistent. Before comparing candidates at all, we fixed threads, KV cache, and checkpoints, and along the way tested two forms of speculative decoding that looked promising in early sweeps and fell apart under load.

Development result

Selected runtime configurations

Development

Docker baseline: 5.3 tokens per second and 4.77 GB, still rising. Resource caps: 6.72 tokens per second and 4.44 GB. Native default: 29.78 tokens per second and 3,519 MiB physical footprint. Six threads with unified KV: 31.09 tokens per second and 3,137 MiB. Draft speculation: 24.72 tokens per second; host memory was not reliably measured.

Fig. 7 Our runtime sequence moves from the Docker baseline through resource caps, native execution, thread tuning, and speculation. These experiments guide configuration; they never rank submission models.

Rejected

Draft-model speculation

Acceptance reached 98.4%, but generation still fell from 30.84 to 24.72 tok/s.

Rejected

N-gram speculation

Acceptance ran only 12–22%, so lookup and verification overhead outweighed the gain.

Adopted

Six-thread cap

Raising the thread count to ten cut decode to about 4.4 tok/s. We had no temperature reading to explain why.

Adopted

Unified KV, two checkpoints

This configuration cut retained state and sped up prefill without slowing decode.Unified KV shares capacity across active slots instead of reserving a fixed block per slot.

Six threads and unified KV got us to 31.09 tok/s at 3,137 MiB, 83% of the estimated weight-bandwidth ceiling, on the development host. None of this configuration travels with the submitted GGUF; it's the fifth method from our list, the runtime work that stays ours to keep but never enters the score. It bought us trustworthy numbers, and that finally let us run the first real model-versus-model scoreboard.

09

Six smaller questions

What we tried and didn't need

Six more levers, each tested once we had a trustworthy runtime to test them on. None of them changed the file we submit, but each one closed a question we'd otherwise still be asking.

GGUF-contained

Mixed tensor or layer quantization

Uniform quantization treats every tensor the same. Mixed precision keeps the parts most sensitive to rounding — embeddings, the output head, the final blocks — at a higher bit width while compressing the rest, and only pays off if the recovered accuracy is worth the extra bytes and the slower kernel those tensors fall back to. We tried it twice, on two different models: on the Qwen3 1.7B ladder, a Q3_K_M body with a Q6_K head against uniform Q3_K_M, and an IQ4_XS variant with the same higher-precision head against plain IQ4_XS; later, on Math-Expert, a Q4_0 body with a Q6_K or Q8_0 tied embedding, and separately a Q4_0 body with Q5_0 in the last four blocks.

None of it held up. The Qwen3 head swap fell to 66% ARC-Easy, worse than uniform Q3_K_M, and the IQ4_XS variant came out both slower and larger than plain IQ4_XS. Math-Expert's mixed layouts recovered no meaningful accuracy either. We rejected every mixed layout we tested — uniform quantization, chosen for the kernel it runs on rather than its nominal precision, kept winning.

GGUF-contained if supported

Structured and unstructured pruning

Structured pruning removes whole layers or low-rank factors a dense runtime can actually skip; unstructured pruning zeros individual weights, and only helps if the file format and kernel give credit for the zeros. We tested both: a Qwen control removed one layer and repeated throughput and ARC-Easy, while the BitCPM branch — covered in full in the vocabulary-pruning chapter above — measured singular-value reconstruction error on its large matrices and checked whether dense GGUF storage could benefit from sparse zeros.

Removing one Qwen layer gained about 3.7% decode speed but cost two ARC-Easy points, worth a full accuracy-score point under the scoring formula — a bad trade for a few percent of throughput. BitCPM's ternary matrices stayed full-rank, with rank-2048 factorisation error near 0.80 before quantization, and dense GGUF storage gives unstructured zeros no credit at all. We rejected every pruning branch we tested; the one pruning result we kept was vocabulary pruning, which prunes tokens rather than weights.

GGUF-contained

Smaller architectures

A smaller model moves fewer weight bytes per decoded token and usually costs less RSS — it only wins if the capability it gives up is smaller than what performance and efficiency pay back. That question isn't a side branch here; it's most of the story on this page. The four-task scale study in the second chapter, the specialist search, and the second widening further down are all, at bottom, this same question asked of a different candidate set, and all three point the same way: the 0.6B–1.7B region is where both roads, and every finalist this story ends with, actually live. See those sections for the full results.

GGUF-contained after training

Distillation and fine-tuning

Distillation transfers a larger teacher's behaviour into a smaller model; fine-tuning adapts a base model to a task — here, mathematical reasoning. Either can raise accuracy without growing the file, and Math-Expert, the specialist search's raw leader, is exactly this kind of result: a Qwen3-0.6B fine-tune on OpenMathReasoning-mini. We screened several more finished public derivatives in that same search (its provenance section records why each one that didn't make the cut fell short), but ran no new training of our own.

We kept Math-Expert as a finished, tested fine-tune, and deferred new training rather than submit something we couldn't fully validate: a GPU, a full dataset audit, and a reproducible conversion path were all missing pieces we couldn't fill in the time we had.

Stored layout plus engine-only policy

Tensor layout, alignment, and runtime repacking

GGUF fixes how tensors are stored on disk, but the engine can still repack them into a faster in-memory layout at load time. Repacking can speed up a kernel, but it costs the RSS budget to do it. A 4B development control compared the default runtime with repacking disabled; we also reviewed a custom alignment or packing scheme on top of that, and decided against building one.

Disabling repacking cut the tested 4B footprint from about 3,236 to 602 MiB, a real product win, without a clear speed loss. But an unsupported layout that fails to load doesn't get a partial score, it gets zero, so we kept the stock GGUF layout for compatibility rather than risk it. No-repack stays a product-only setting — it's method ten on our list, the boundary this whole story keeps running into.

Engine-only

Context size and KV-cache implications

Context length sets how many token states the KV cache holds. A shorter or quantized cache can save memory, but it doesn't touch the weight traffic a batch-one decode has to move every token regardless of context. We'd already replaced runtime defaults with explicit context and KV limits back in the runtime-tuning chapter, which stopped memory growth and made later comparisons trustworthy — and the participant profiler fixes its own workload anyway, at a 512-token prompt and 128 generated tokens.

No GGUF-level context metadata we could set would change what gets scored, so we rejected context metadata as a profiler-facing lever. We keep the bounded context and KV settings we already adopted, but only as product configuration.

10

Checkpoint

The first scoreboard: Qwen3 1.7B wins, narrowly

With quantization, deduplication, and the runtime settled, we ran six candidates through the actual participant profiler. Qwen3 1.7B Q4_0, pure and tied, won at 72.4653, ahead of Qwen3.5 0.8B by 0.92 points and comfortably ahead of both the BitCPM ternary bet and a larger Qwen3.5 4B candidate.

Official profiler result

Score by component

Direct
Fig. 8 Contributions use the executable profiler's fixed 15 tok/s reference. ARC-Easy supplies the accuracy proxy; we still don't have the judging-panel score.

Official profiler result

Speed against memory

Direct
Fig. 9 We plot generation throughput against peak RSS. Higher throughput and lower RSS both raise the total. Circle area represents the ARC-Easy proxy.
Scalar measurements
Modeltok/sFirst tokenPeak RSSARC-EasyTotal
Math-Expert 0.6B Q4_K_MScalar leader · 396.7 MB12.7223.61 s540 MiB68% 54.2–79.277.9324
Qwen3.5 0.8B Q4_0Risk-adjusted recommendation · 507.2 MB12.6316.62 s670 MiB64% 50.1–75.975.3895
Qwen3 1.7B Q4_0Previous scalar leader · 974.2 MB9.7935.37 s1,116 MiB72% 58.3–82.572.4653
Qwen3.5 0.8B Q4_K_M9.7428.26 s695 MiB68% 54.2–79.271.54
BitCPM4 8B TQ2_00.81584.15 s2,307 MiB88% 76.2–94.459.18
Qwen3.5 4B IQ4_XS1.13395.19 s2,627 MiB76% 62.6–85.752.93

The final Qwen model has the shortest measured first-token latency, 16.62 seconds against Math-Expert's 23.61. Neither number enters the score.

The table above already shows Math-Expert and the final Qwen3.5 0.8B ahead of the 1.7B model that won this first round. That's the ledger as it stands today: two later rounds moved the winner twice more. At the time of this first scoreboard, though, 1.7B genuinely led, and we treated it as settled long enough to ask a harder question: would the public webpage's cohort-relative scoring formula, instead of the fixed 15, change anything?

We checked, and the full breakdown, at every cohort floor we tested, is in the operational appendix under Website-relative sensitivity. It's a separate check, never blended with the profiler result above.

11

Method six

Finding the real CPU, and watching the ranking flip

Everything so far assumed the audit binary runs scalar only, because that's how the executable profiler we could get our hands on was built. But the competition's own public page describes hardware that plausibly supports the vector configuration defined above. We couldn't be sure which one the final judging actually runs, so we built both and measured the same seven models on each.

AVXON AVX2ON FMAON F16CON NATIVEOFF AVX-512OFF

Matched models and scoring function

Total score by CPU method

ScalarVector

Loading paired scalar and vector score evidence.

Fig. 10 Each model carries a scalar and a vector fixed-15 total, both using its ARC-Easy-50 proxy. The finalist scalar rows are participant-profiler results; the quantization-ladder rows are controlled scalar screens. Outlined bars mark the highest total in each CPU configuration.

Turning the vector configuration on reordered the field as well as sped it up. Q4_K_M, penalised by the generic-C fallback under quantization's first lesson, jumped ahead of pure Q4_0 once it had a real vector kernel to run on: 80.4484 against 80.2818, a lead earned entirely by 59.7 MiB less repacking RSS, not by anything about the model, and not by any change to its measured accuracy. Even BitCPM, road two's abandoned bet, sped up 9.235× and became operationally viable, though its total still couldn't catch the Qwen variants.

Scalar and vector measurements
Scalar and vector benchmark results for seven GGUF models across the quantization and expanded model searches.
ModelScalar → vector pp512Scalar → vector tg128Decode gainEst. profiler RSSARC-Easy proxyTotal
Math-Expert 0.6B Q4_K_MExpanded-search leader21.7268 → 153.935112.6339 → 39.23203.105×759.7 MiB+202.9 MiB68%81.8803
Qwen3.5 0.8B Q4_0Risk-adjusted recommendation; tensor-identical vector source31.0781 → 98.009412.6955 → 27.15092.139×928.1 MiB+240.8 MiB64%79.4104
Qwen3 1.7B Q4_0Previous scalar choice14.7166 → 47.07169.9869 → 16.89271.691×2,049.4 MiB+916.3 MiB72%80.2818
Qwen3 1.7B Q4_K_MQuantization-ladder leader7.1859 → 55.45545.2954 → 15.67142.959×1,989.7 MiB+806.2 MiB72%80.4484
Qwen3 1.7B Q5_K_M6.5613 → 24.32314.7839 → 12.71912.659×1,364.6 MiB+0.1 MiB76%79.6307
Qwen3 1.7B IQ4_XS3.2063 → 23.93642.4961 → 14.06445.635×1,082.3 MiB+0.5 MiB70%80.1089
BitCPM4 8B TQ2_00.8762 → 13.65690.8108 → 7.48769.235×2,316.4 MiB+0.1 MiB88%72.5121
14

Method eight

Teaching the file to behave, since nothing else is in the room

Every method so far shrinks bytes or speeds up decode. This one changes what the model says: judges chat with the bare GGUF directly, no system prompt and no application layer standing between them and the model, so the tutoring persona has to travel inside the file or it doesn't exist during evaluation.

Metadata-only derivation of the previously packaged Qwen3.5 GGUF Qwen3.5 source quant Qwen3.5 0.8B Q4_0 507.2 MB Embed tutor policy metadata and ChatML only Embedded tutor policy English · thinking off sampling defaults write final GGUF no tensor conversion Evaluated model Muta Tutor Qwen3.5 507.2 MB +1.5 KB metadata Tensor-identity verification 320 model tensors compared all tensor payloads identical Only the evaluated GGUF is submitted. Runtime flags, engine patches, and application code are not included.

The Qwen3.5 0.8B Q4_0 source is 507.2 MB. The build adds the tutor policy, English language tag, non-thinking ChatML template, and sampling defaults without changing model tensors. Verification confirms that all 320 model tensors are identical.

Fig. 16 The previous packaged Qwen3.5 model differed from its source quant only in GGUF metadata. The tuned finalist must receive and revalidate the same policy before submission.

We replace the chat template with a clean ChatML template that injects the Muta tutoring persona as the system turn whenever the client sends none, and always opens the assistant turn with an empty <think></think> block, forcing a direct answer instead of a long reasoning trace. That last part isn't cosmetic. A live four-prompt acceptance test with unrestricted thinking left both finalists spending their entire 256-token allowance on hidden reasoning and returning no usable answer at all, a minute of silence per prompt at audit-box speed. Disabling thinking fixed that. Sampling defaults (temperature 0.4, top-p 0.9, min-p 0.05, repeat penalty 1.05) are read straight from the file by llama-server, and the persona runs about 130 tokens, so a judge's first turn pays little extra prefill for it.

We verified this on llama.cpp's Jinja engine and llama-cpp-python's separate jinja2 path. The previous Qwen3.5 package injected the persona and stopped correctly, but the hardest proof prompt still failed. The fine-tuned finalists below have not yet repeated this battery; measured accuracy alone does not clear that requirement.

15

Method nine

Fine-tuning the two finalists

We trained 15 candidates across Qwen3.5 0.8B and Qwen2.5 1.5B, varying LoRA rank, BF16 LoRA versus QLoRA, data mixture, learning rate, and training length. Every survivor was exported to its deployment quantization and evaluated as a GGUF; training loss did not select a winner.

The first eight runs did not improve ARC-Easy-500. Most training examples were long solutions or tutor dialogue, while the profiler scores short answer continuations. The original data builder also omitted the leading space between Answer: and the candidate continuation, changing the BPE tokens used during training. We corrected both problems, filtered against held-out questions, and used only source training splits. The final Qwen2.5 mixture excludes OpenBookQA because its published licence is unclear.

Initial sweep

Eight runs · no promotion

Balanced and reasoning-heavy mixtures improved validation loss but did not improve the exact profiler task.

Metric-aligned sweep

Seven runs · two promotions

Raw multiple-choice continuations, corrected token boundaries, verified answers, and leakage checks improved both finalists.

Matched 500-item evaluation

Accuracy before and after tuning

ControlFine-tuned

Loading fine-tuning accuracy results.

Fig. 17 Qwen3.5 improved from 55.2% to 70.2%; Qwen2.5 improved from 74.4% to 77.8%. Both comparisons use matched exports and the same 500-item metric.

Fine-tuned GGUFs

Total score by CPU configuration

ScalarVector

Loading fine-tuned finalist scores.

Fig. 18 The scalar and vector rankings remain split: tuned Qwen3.5 leads scalar; tuned Qwen2.5 leads vector.
Matched control, performance, and memory results

Qwen3.5 gained 15.0 accuracy points and approximately 7.5 total points in either CPU configuration. Qwen2.5 gained 3.4 accuracy points and approximately 1.7 total points. Throughput and memory were effectively unchanged within each matched pair, so these score gains come from accuracy rather than a deployment trade-off.

Secondary held-out checks

16

Current decision

Fine-tuning raises both scores; the runtime split remains

The matched 500-item comparison still selects one model per CPU configuration. Fine-tuning raises both totals without materially changing speed or memory.

Scalar configuration, n=500

Fine-tuned Qwen3.5 0.8B · 80.3664

70.2% ARC-Easy, 13.60 tok/s, and an estimated 691 MiB profiler RSS.

Vector configuration, n=500

Fine-tuned Qwen2.5 1.5B · 84.1387

77.8% ARC-Easy, 17.44 tok/s, and an estimated 1,706 MiB profiler RSS.

Quantization decides the available CPU kernel; fine-tuning only changes accuracy on top of it. So the final choice still comes down to which CPU configuration the evaluator actually runs.

17

The receipts

Every experiment, adopted, rejected, and deferred

The measurements above produced the decisions below. Adopted changes appear first; use the filters to see rejected, neutral, and deferred tests.

18

Challenge requirements

Requirement status

Model selection also has to satisfy the published challenge requirements. Each entry pairs one requirement with what we have actually implemented.

Challenge requirements follow the Africa Deep Tech Challenge 2026 page.

A

Operational appendix

Profiler results and comparison tables

The underlying measurements, live profiling controls, stored results, and mixed-machine archive all stay available below.

Profiler metadata

Loading metadata…

Official profiler resultDirect GGUF campaignLoading campaign evidence…

Profiler-parity estimateCandidate screening with the scalar audit kernelLoading reconstructed profiler evidence…

Controlled vector proxyPaired scalar and vector measurements, seven-model ladderLoading the seven-model instruction-set comparison…

Controlled vector proxyPaired scalar and vector measurements, eight-model architecture screenLoading the eight-model architecture screen…

Website-relative sensitivityVector measurements with the cohort-relative formulaLoading sensitivity data…

Mixed-machine archiveHistorical profiler archiveDifferent machines and engine configurations; excluded from the campaign ranking.

Historical display only: S_total = 0.50·S_acc + 0.30·S_perf + 0.20·S_eff − 10·thermal · legacy S_perf = min(TPS / TPS_max, 1) × 100 · TPS_max = fastest stored run · S_eff = (7 GB − peak RAM) / 7 GB × 100 · crash or OOM ⇒ disqualified (0).

Stored model runs

Muta IQ · Offline adaptive tutoring for maths and scientific reasoning.

We keep the full methods and raw measurements in the repository benchmark archive.

Continue to Gate 2