can-local-ai-help-me-cheat-on-my-math-test.mdxPost FileCan local AI help me cheat on my math test?August 25, 2026 by discostewChecking vision capable models against college level math problems. Testing both their vision capability and reasoning ability. Correct answers first. Then total wall time, output tokens, and measured generation rate on the same five native-image problems.local AIvision modelsmath benchmarkApple SiliconFastest 5/5 resultGemma 4 12B 8-bit · 8m 21.4sQwen3.8 quality setting27B Q4 · xhigh · 5/5 · 21,230 tokensThe exact five test imagesThese are the actual PNGs supplied to every native-vision route below.01 Tilted square and circlechallenge-packages/geometry-area-challenges-v1/images/01-tilted-square-circle.png02 Overlapping circleschallenge-packages/geometry-area-challenges-v1/images/06-overlapping-circles.png03 Nested square and circlechallenge-packages/geometry-area-challenges-hard-v1/images/01-nested-square-circle-square.png04 Quarter-circle lenschallenge-packages/geometry-area-challenges-hard-v1/images/04-opposite-corner-quarter-circle-lens.png05 Tilted square areascreenshots/source/tilted-square-area.pngThe resultsThirteen recorded local runs, ranked by correct answers and then summed task wall time. Model load and unload are excluded.1Gemma 4 12BDense8-bitThinking onCorrect5/5Wall time8m 21.4sTotal tokens20,305Generation41.21 tok/s2Qwen3.8 27BDense4-bitxhigh thinkingCorrect5/5Wall time10m 33.1sTotal tokens21,230Generation34.94 tok/s3Qwen3.5 35B-A3BMoE4-bitThinking onCorrect5/5Wall time11m 9.4sTotal tokens56,443Generation85.83 tok/s4Qwen3.5 35B-A3BMoE8-bitThinking onCorrect5/5Wall time17m 7.3sTotal tokens70,264Generation69.22 tok/s5Qwen3.5 9BDense8-bitThinking onCorrect5/5Wall time19m 19.0sTotal tokens74,240Generation64.84 tok/s6Qwen3.8 27BDenseoQ8xhigh thinkingCorrect5/5Wall time21m 39.3sTotal tokens27,325Generation21.46 tok/s7Gemma 4 26B-A4BMoE8-bitThinking onCorrect4/5Wall time3m 5.2sTotal tokens13,894Generation78.65 tok/s8Gemma 4 26B-A4BMoEQAT Q4 base / Q8 MoEThinking onCorrect4/5Wall time5m 50.0sTotal tokens28,921Generation84.68 tok/s9Qwen3.8 27BDense4-bitlow thinkingCorrect4/5Wall time7m 12.3sTotal tokens14,694Generation36.09 tok/s10Muse Glimmer 30BMoE4-bitThinking onCorrect4/5Wall time15m 1.6sTotal tokens28,563Generation32.86 tok/s11Muse Glimmer 30BMoE8-bitThinking onCorrect4/5Wall time17m 24.6sTotal tokens20,268Generation20.08 tok/s12Qwen3.6 27BDense4-bitDefault thinkingCorrect4/5Wall time29m 12.2sTotal tokens60,679Generation35.13 tok/s13Qwen3.8 27BDense4-bitmedium thinkingCorrect3/5Wall time8m 51.8sTotal tokens18,316Generation36.13 tok/sTest machineEvery retained result on this page ran locally on the same Mac Studio.ModelMac Studio (Mac15,14)ChipApple M3 UltraCPU32-core CPU (24 performance + 8 efficiency)GPU80-core GPUMemory512 GB unified memoryOSmacOS 26.6.1 (25G76)What these five problems sayQwen3.8’s reasoning is the real upgrade here.At the same Q4 size, Qwen3.8 xhigh went 5/5 with 21,230 output tokens. Qwen3.6 went 4/5 with 60,679—2.86× as many—and took 29m 12.2s. That is consistent with much stronger 3.8 reasoning/post-training, although this benchmark cannot separate base training from post-training.Qwen3.8 needs xhigh for this quality bar.xhigh: 5/5, 21,230 tokens. Medium: 3/5, 18,316. Low: 4/5, 14,694. Low is cheaper, but it lost an answer; medium lost two.Q4 is the obvious Qwen MoE choice.Qwen3.5 35B-A3B kept 5/5 at both quantizations. Q4 used 56,443 tokens in 11m 9.4s versus Q8’s 70,264 in 17m 7.3s.The Gemma MoE result is more nuanced.Its QAT Q4/Q8 hybrid matched full Q8 at 4/5, so quality held, but one long wrong trace pushed wall time to 5m 50.0s versus 3m 5.2s at Q8. Choose the smaller load unless latency is the priority.Do not spend memory on Qwen3.8 oQ8 for this workload.Q4 xhigh and oQ8 both scored 5/5. Q4 used 21,230 tokens in 10m 33.1s; oQ8 used 27,325 in 21m 39.3s.Best tested quantization by modelThe practical choice from this retained five-problem evidence, not a claim about every workload.8-bitGemma 4 12B · Dense5/5 · 8m 21.4s · 20,305 tokensThe retained passing choice: keep the small dense model at Q8 unless a Q4 run proves parity.4-bit · xhighQwen3.8 27B · Dense5/5 · 10m 33.1s · 21,230 tokensThe quality choice: 5/5. Medium fell to 3/5 and low to 4/5; oQ8 kept 5/5 but took 21m 39.3s.4-bitQwen3.5 35B-A3B · MoE5/5 · 11m 9.4s · 56,443 tokensKept the same 5/5 as Q8 with 56,443 rather than 70,264 output tokens and a shorter wall time.8-bitQwen3.5 9B · Dense5/5 · 19m 19.0s · 74,240 tokensThe retained passing choice for this small dense model; it reached 5/5.QAT Q4 base / Q8 MoEGemma 4 26B-A4B · MoE4/5 · 5m 50.0s · 28,921 tokensSame 4/5 score as full Q8 with the smaller load; Q8 cut wall time to 3m 5.2s in this run, so use it only when latency matters more than memory.4-bitMuse Glimmer 30B · MoE4/5 · 15m 1.6s · 28,563 tokensMatched Q8 at 4/5 while reducing wall time from 17m 24.6s to 15m 1.6s.4-bit reference onlyQwen3.6 27B · Dense4/5 · 29m 12.2s · 60,679 tokensUseful as the before case, but it missed one problem and generated far more tokens than Qwen3.8 Q4 xhigh.Five-problem native-image reasoning test · Local retained results
2Qwen3.8 27BDense4-bitxhigh thinkingCorrect5/5Wall time10m 33.1sTotal tokens21,230Generation34.94 tok/s
3Qwen3.5 35B-A3BMoE4-bitThinking onCorrect5/5Wall time11m 9.4sTotal tokens56,443Generation85.83 tok/s
4Qwen3.5 35B-A3BMoE8-bitThinking onCorrect5/5Wall time17m 7.3sTotal tokens70,264Generation69.22 tok/s
6Qwen3.8 27BDenseoQ8xhigh thinkingCorrect5/5Wall time21m 39.3sTotal tokens27,325Generation21.46 tok/s
8Gemma 4 26B-A4BMoEQAT Q4 base / Q8 MoEThinking onCorrect4/5Wall time5m 50.0sTotal tokens28,921Generation84.68 tok/s
10Muse Glimmer 30BMoE4-bitThinking onCorrect4/5Wall time15m 1.6sTotal tokens28,563Generation32.86 tok/s
11Muse Glimmer 30BMoE8-bitThinking onCorrect4/5Wall time17m 24.6sTotal tokens20,268Generation20.08 tok/s
12Qwen3.6 27BDense4-bitDefault thinkingCorrect4/5Wall time29m 12.2sTotal tokens60,679Generation35.13 tok/s
13Qwen3.8 27BDense4-bitmedium thinkingCorrect3/5Wall time8m 51.8sTotal tokens18,316Generation36.13 tok/s
8-bitGemma 4 12B · Dense5/5 · 8m 21.4s · 20,305 tokensThe retained passing choice: keep the small dense model at Q8 unless a Q4 run proves parity.
4-bit · xhighQwen3.8 27B · Dense5/5 · 10m 33.1s · 21,230 tokensThe quality choice: 5/5. Medium fell to 3/5 and low to 4/5; oQ8 kept 5/5 but took 21m 39.3s.
4-bitQwen3.5 35B-A3B · MoE5/5 · 11m 9.4s · 56,443 tokensKept the same 5/5 as Q8 with 56,443 rather than 70,264 output tokens and a shorter wall time.
8-bitQwen3.5 9B · Dense5/5 · 19m 19.0s · 74,240 tokensThe retained passing choice for this small dense model; it reached 5/5.
QAT Q4 base / Q8 MoEGemma 4 26B-A4B · MoE4/5 · 5m 50.0s · 28,921 tokensSame 4/5 score as full Q8 with the smaller load; Q8 cut wall time to 3m 5.2s in this run, so use it only when latency matters more than memory.
4-bitMuse Glimmer 30B · MoE4/5 · 15m 1.6s · 28,563 tokensMatched Q8 at 4/5 while reducing wall time from 17m 24.6s to 15m 1.6s.
4-bit reference onlyQwen3.6 27B · Dense4/5 · 29m 12.2s · 60,679 tokensUseful as the before case, but it missed one problem and generated far more tokens than Qwen3.8 Q4 xhigh.